Skip to content

[bench] sparse bitpacked filter and take - #9900

Open
robert3005 wants to merge 2 commits into
developfrom
rk/fastlanes-sparse-benchmarks
Open

robert3005 wants to merge 2 commits into
developfrom
rk/fastlanes-sparse-benchmarks

Conversation

@robert3005

@robert3005 robert3005 commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Adds sparse bitpacked filter and take benchmarks for u8, u16, u32, and u64, providing a baseline for the extraction optimization in #9723.

@robert3005
robert3005 added this pull request to stack #9901 September 16, 2026 02:46
@robert3005 robert3005 changed the title perf(fastlanes): Benchmark sparse extraction [bench] sparse bitpacked filter and take Sep 16, 2026
@robert3005 robert3005 added the changelog/chore A trivial change label Sep 16, 2026
@codspeed

codspeed Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Merging this PR will improve performance by 34.45%

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ 2 benchmarks measured no execution time

Nothing ran under measurement, usually because the compiler removed the code under test. These results are not comparable, so they count as unchanged.

Preventing compiler optimizations

⚠️ 12 benchmarks spent significant time in system calls

System calls cannot be consistently instrumented, so they are not included in the measure, which understates the real cost. Please switch to the Walltime instrument to accurately measure system calls.

Measurement and system calls

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 2 improved benchmarks
✅ 2181 untouched benchmarks
🆕 18 new benchmarks
⏩ 359 skipped benchmarks1

Performance Changes

Mode Benchmark BASE HEAD Efficiency
⚡ ⚠️ Simulation density_sweep_dense_runs[0.001] 48.2 µs 30.1 µs +60.13%
⚡ WallTime bitpack_blocked_compress_avx2 7.6 µs 6.7 µs +12.89%
🆕 WallTime filter_avx512[64] N/A 2.2 µs N/A
🆕 WallTime filter_avx512[8] N/A 1.1 µs N/A
🆕 WallTime filter_avx512[80] N/A 2.3 µs N/A
🆕 WallTime threshold_avx512[64] N/A 2.9 µs N/A
🆕 WallTime threshold_avx512[8] N/A 1.8 µs N/A
🆕 WallTime threshold_avx512[80] N/A 3.2 µs N/A
🆕 WallTime filter_avx2[64] N/A 2.4 µs N/A
🆕 WallTime filter_avx2[8] N/A 1.1 µs N/A
🆕 WallTime filter_avx2[80] N/A 2.5 µs N/A
🆕 WallTime threshold_avx2[64] N/A 3 µs N/A
🆕 WallTime threshold_avx2[8] N/A 1.9 µs N/A
🆕 WallTime threshold_avx2[80] N/A 3.3 µs N/A
🆕 WallTime filter_neon[64] N/A 2.9 µs N/A
🆕 WallTime filter_neon[8] N/A 1.1 µs N/A
🆕 WallTime filter_neon[80] N/A 3.1 µs N/A
🆕 WallTime threshold_neon[64] N/A 4.5 µs N/A
🆕 WallTime threshold_neon[8] N/A 3.1 µs N/A
🆕 WallTime threshold_neon[80] N/A 4.8 µs N/A

Tip

Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.


Comparing rk/fastlanes-sparse-benchmarks (6cd3ac0) with develop (1ee4e3a)

Open in CodSpeed

Footnotes

  1. 359 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

This PR has been marked as stale because it has been open for 14 days with no activity. Please comment or remove the stale label if you wish to keep it active, otherwise it will be closed in 7 days

@github-actions github-actions Bot added the stale This PR is stale and will be auto-closed soon label Oct 3, 2026
@robert3005 robert3005 removed the stale This PR is stale and will be auto-closed soon label Oct 3, 2026
Signed-off-by: Will Manning <will@willmanning.io>
Signed-off-by: "Robert Kruszewski" <github@robertk.io>
@robert3005
robert3005 force-pushed the rk/fastlanes-sparse-benchmarks branch from f17daff to 6078880 Compare October 8, 2026 21:36
@joseph-isaacs

Copy link
Copy Markdown
Contributor

Do we need 72 benchmarks?

Signed-off-by: Robert Kruszewski <github@robertk.io>
@robert3005

Copy link
Copy Markdown
Contributor Author

you're correct that we don't

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/chore A trivial change

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants